Back

DNA Research

Oxford University Press (OUP)

Preprints posted in the last 90 days, ranked by how well they match DNA Research's content profile, based on 26 papers previously published here. The average preprint has a 0.02% match score for this journal, so anything above that is already an above-average fit.

1
A highly contiguous genome assembly of Cyclamen persicum to accelerate functional genomics and breeding

Shirasawa, K.; Akita, Y.; Mizunoe, Y.; Takamura, T.

2026-07-21 genomics 10.64898/2026.07.17.739093 medRxiv
Top 0.1%
56.0%
Show abstract

Cyclamen is an economically important ornamental plant widely cultivated for its diverse floral characteristics and adaptation to cool climates. Despite its horticultural significance, genomic resources for this species remain limited, hindering molecular studies and genomics-assisted breeding. Here, we report the first highly contiguous nuclear genome assembly of C. persicum generated using high-fidelity long-read sequencing. The assembled genome spans 1.48 Gb, consisting of 126 contigs with an N50 length of 52.3 Mb. Telomeric repeat analysis identified eight contigs containing telomeric sequences at both ends, suggesting the presence of near-complete chromosome assemblies. Genome completeness assessment using BUSCO indicated 98.1% completeness. Repetitive sequences occupied 82.9% of the assembly, with long terminal repeat retrotransposons accounting for 42.1% of the genome. A total of 40,223 protein-coding genes were predicted, with a complete BUSCO score of 95.7%. Comparative orthogroup analysis with five representative eudicot species identified 430 orthogroups specific to C. persicum and 363 orthogroups shared exclusively between C. persicum and Primula kwangtungensis, indicating the presence of both lineage-specific and Primulaceae-conserved gene families. These findings provide critical insights into gene family evolution within Primulaceae and establish an essential comparative framework for future genomic studies. The genome resource presented here provides an invaluable foundation for investigating genome evolution, gene function, and trait-associated loci in cyclamen, effectively facilitating molecular breeding and genetic improvement in this ornamental species.

2
Chromosome-level genome assembly of the Northeast China Brown Frog (Rana dybowskii)

zhang, y.; Wang, D.; Zhao, R.; Li, S.; Zheng, X.; Hu, G.

2026-06-15 genomics 10.64898/2026.06.11.731602 medRxiv
Top 0.1%
46.2%
Show abstract

Rana dybowskii is distributed across Northeast Asian and represents a valuable medical resource. A high-quality assembly of the genome has not yet been reproted. This species has 2n=24 chromosomes, but a huge genome size that estimated at 3.5 ~4.6 Gb in the previous studies. The relatively large chromosome size, exceeding hundreds of megabases, may result in difficulties of obtaining a complete chromosome level genome. Here, we constructed a chromosome-level genome assembly of R. dybowskii by integrating PacBio HiFi long-read sequencing for de novo assembly and CiFi (3C coupled with HiFi sequencing) for scaffolding. The final assembly consists of 12 chromosomes with a total of 3.95 Gb and a scaffold N50 length of 455 Mb. BUSCO assessment using the tetrapoda_odb12 database identified 94.2% complete and 0.5% fragmented orthologs, suggesting a high level of completeness of the assembly. Genomic annotation revealed that repetitive sequences comprise over 53% of the assembly, with retroelements and DNA transposons accounting for 22% and 25%, respectively. A total of 43,999 protein-coding genes were predicted with the assistance of RNA-seq reads from four tissues (muscle, eye, testis and skin). This high-quality chromosome-level reference genome provides a valuable genomic resource for advancing genetic studies of the species.

3
Chromosome-level genome assembly of Calotes wangi with dynamic colour variation

Qiu, X.; Wang, Y.; Wen, J.; Chen, Y.; Zhao, L.; Jian, J.; Yang, W.

2026-07-10 evolutionary biology 10.64898/2026.07.07.736949 medRxiv
Top 0.1%
30.6%
Show abstract

The Wangs garden lizard, Calotes wangi, is a widely distributed agamid species in Southern China and Northern Vietnam and exhibits pronounced colour variation and rapid body colour change. Despite increasing interest in the genomic basis of colour variation, chromosome-level genomic resources remain limited in agamid lizards. Here, we generated a chromosome-level reference genome of C. wangi using PacBio HiFi sequencing and Hi-C scaffolding. The final genome assembly was approximately 1.66 Gb in size and comprised 6 macrochromosomes and 11 microchromosomes, with a contig N50 of 110.09 Mb and 98.9% complete BUSCO genes. A total of 20,442 protein-coding genes were annotated. Comparative genomic analyses identified 297 significantly expanded gene families, with enriched functions associated with steroid metabolism, chromatin regulation, and epigenetic processes. This high-quality genome assembly provides an important genomic resource for future studies of colour variation, phenotypic plasticity, and evolutionary diversification in agamid lizards.

4
Chromosome-scale genome assembly of Cycas revoluta provides insights into cycad diversification and species diversity

Sato, M. P.; Aoyagi, Y. B.; Yoshitake, K.; Toyama, Y.; Iimura, H.; Matsuoka, F.; Toyoda, A.; Akagi, T.; Okuda, S.; Ito-Inaba, Y.; Shirasawa, K.

2026-07-17 genomics 10.64898/2026.07.11.737888 medRxiv
Top 0.1%
29.3%
Show abstract

Cycads are an ancient group of seed plants. Despite their ancient origin, many extant cycad genera exhibit high species diversity. The specialized reproductive traits in cycads, dioecy governed by the XY sex-determination system and the elaborate co-evolutionary synergy with insect pollinators, may facilitate lineage diversification. Here, we present a chromosome-scale genome sequence of Cycas revoluta, the species in which plant spermatozoids were discovered in 1896. The genome sequence spanned 11.6 Gb, 98.4% of which were anchored onto the 11 cycad chromosomes. Repetitive sequences occupied 9.8 Gb, and 31,481 genes were predicted. Based on this genome assembly, the X- and Y-associated genomic regions were characterized, and candidate genes for sex determination were identified. In addition, the genomic positions of genes for sex-related traits were determined. Subsequently, we analyzed transcriptomes for thermogenesis responsible for attracting insect pollinators. Through these comprehensive analyses, we provide new insights into the genomic basis of cycad diversification and species diversity.

5
A high-quality chromosome-scale reference genome assembly for Asparagus racemosus var. CIM-Shakti (Shatavari), a medicinal plant of Ayurvedic importance

Tyagi, S.; Sharma, A.; Shivani, K.; Gupta, V.; Paterson, A. H.; Trivedi, P. K.

2026-06-11 bioinformatics 10.64898/2026.06.07.730773 medRxiv
Top 0.1%
12.3%
Show abstract

Asparagus racemosus Wild., commonly known as Shatavari, is an important medicinal plant in Ayurveda and is valued for its steroidal saponins, particularly shatavarin compounds, which contribute to its adaptogenic, galactagogue, immunomodulatory, and therapeutic properties. Despite its medicinal and economic importance, genomic resources for this species have remained limited, restricting molecular breeding, pathway discovery, and comparative evolutionary studies within Asparagaceae. Here, we report a high quality chromosome scale reference genome assembly of A. racemosus var. CIM Shakti generated using PacBio HiFi long read sequencing and Omni C chromatin conformation scaffolding. The pseudo haploid assembly spans 817 Mb across 53 scaffolds, with a scaffold N50 of 98.50 Mb, L50 of 5, and a largest scaffold of 113.80 Mb. Ten major chromosome scale pseudomolecules were resolved, corresponding to the haploid chromosome complement of A. racemosus. The assembly showed high gene space completeness, with BUSCO completeness of 99.8% against the Eukaryota dataset and 98.0% against the Embryophyta dataset. BlobToolKit profiling further supported assembly quality, with GC content of approximately 39 to 40% and no major evidence of contamination. EDTA based repeat annotation identified 580.93 Mb of interspersed repetitive elements, accounting for 71.06% of the 817.57 Mb genome assembly. The repeat landscape was dominated by LTR retrotransposons, particularly Gypsy elements, which accounted for 25.01% of the assembly, followed by unclassified LTR elements at 26.58% and Copia elements at 4.84%. Structural and functional annotation identified 29,199 protein coding genes represented by 29,199 transcript models, 138,433 exons, and 125,201 CDS features. The annotation was structurally robust, with an average gene length of 4,605.1 bp, 4.74 exons per transcript, and 97.80% of transcripts containing multiple exons. The CIM Shakti reference genome provides a foundational genomic resource for investigating steroidal saponin biosynthesis, sex chromosome evolution, repeat driven genome expansion, and comparative genomics in Asparagaceae. This assembly will support future studies on medicinal trait improvement, conservation genomics, and genomics assisted breeding of climate resilient Shatavari cultivars.

6
Chromosome-scale genome assembly and annotation of the Vietnamese indica rice cultivar Khang Dan 18

Nguyen, T. Q.; Do, K. H. D.; Vu, T. M.; Hoang, N. V.

2026-08-21 plant biology 10.64898/2026.08.15.742683 medRxiv
Top 0.1%
6.7%
Show abstract

Khang Dan 18 (KD18) is an Oryza sativa L. subsp. indica rice cultivar widely cultivated in northern Vietnam and used as an experimental and breeding background in Vietnamese rice research. Although KD18 has previously been represented in low-depth population resequencing datasets, a contiguous and annotated cultivar-specific genome has not been available. Here, we report a chromosome-scale genome assembly of KD18 generated using Oxford Nanopore long-read and Illumina short-read sequencing. The 395.3-Mb assembly comprises 12 chromosome-scale pseudomolecules containing approximately 95% of the assembled sequence and 99.6% of the predicted protein-coding genes. The assembly showed 97.2% BUSCO completeness, an average Merqury quality value of 46 and a long terminal repeat assembly index of 13.21. A total of 56,546 protein-coding genes representing 71,237 transcripts were predicted, with 99% BUSCO and 98.68% OMArk completeness. These statistics are similar to those of other high-quality genome assemblies that were recently published for different Asian rice cultivars, therefore providing a cultivar-specific genomic resource for research involving KD18 and KD18-derived materials.

7
Phasing of the 'Wonderful' Pomegranate Genome Using Haploid DNA Extracted from Pollen Grains

Lana, G.; Traband, R.; Ferrante, S. P.; Resendiz, M.; Yu, L.; Qu, H.; Eurmsirilerd, E.; Deng, Z.; Roose, M.; Merhaut, D.; Beaulieu, T.; Seymour, D.; Gmitter, F.; Jia, Z.; Chater, J.

2026-07-18 genomics 10.64898/2026.07.13.738218 medRxiv
Top 0.1%
6.6%
Show abstract

The scientific and commercial interest in pomegranate (Punica granatum L.) cultivation has increased noticeably during the last two decades. Because of the high concentration of bioactive compounds and its promising nutraceutical properties, pomegranate has been defined as a functional food. In order to develop advanced genomic tools to improve pomegranate breeding program efficiency, we present the chromosome-scale and haplotype-resolved genome assembly of Wonderful, a pomegranate cultivar widely grown around the world. DNA isolated from diploid leaf tissues was sequenced using long read sequencing technology (PacBio and Nanopore), while DNA extracted from haploid pollen grains was sequenced using a short-reads platform (Illumina). Genomic data from 11 single haploid gamete cells were analyzed using the R package called Hapi to phase the genome. The final genome assembly size was of 372.51 Mbp anchored to eight pairs of homologous chromosomes. The present study provides an insight on the adoption of an innovative and efficient approach for the assembly of haplotype-resolved genomes, which enables a higher resolution of DNA variant detection and offers the opportunity to investigate crossover events in single gamete cells during meiosis.

8
Transcriptomic datasets of melittin- and un-treated murine cervical carcinoma U14 cells

Zhang, R.; Zhang, Y.; Wang, M.; Jiang, J.; Li, Y.; Chen, D.; Yan, T.; Guo, R.

2026-07-28 genomics 10.64898/2026.07.25.739748 medRxiv
Top 0.1%
5.5%
Show abstract

Melittin, a potent amphipathic cationic peptide derived from bee venom, exhibits broad-spectrum antineoplastic efficacy, notably against cervical carcinoma. Despite its established therapeutic potential, the global transcriptional reprogramming orchestrating its acute multi-pathway cytotoxicity remains incompletely understood. To bridge this knowledge gap, we generated the first comprehensive, untargeted RNA-seq dataset profiling the acute phase of melittin-induced cell death in murine cervical carcinoma U14 cells (exposed to 4 g/mL melittin for 20 minutes) alongside untreated controls. Utilizing deep sequencing and rigorous bioinformatics workflows, we quantified genome-wide mRNA abundances and mapped a distinct transcriptomic shift, identifying 254 significantly differentially expressed genes, comprising 158 up- and 96 down-regulated transcripts. Validated by stringent quality control metrics, exceptional genomic mapping rates, and comprehensive functional annotations via the GO and KEGG databases, this high-resolution transcriptomic resource provides a systems-level map of early molecular alterations. All raw and processed sequencing data are publicly available. This transcriptomic resource provides a valuable foundation for elucidating the acute regulatory networks underlying melittin-induced anti-cervical cancer effects. DatasetThe dataset can be accessed through the NGDC website by searching with the BioProject accession number PRJCA068439. Reviewers may use this link for anonymous access during the review process. Direct URL to data: Genome Sequence Archive-CNCB-NGDC

9
Chromosome organization of Entamoeba histolytica and Entamoeba dispar

Kawano-Sugaya, T.; Kobayashi, S.; Kawashima, A.; Saito-Nakano, Y.; Izumiyama, S.; Nozaki, T.; Nakada-Tsukui, K.

2026-07-09 genomics 10.64898/2026.07.06.736064 medRxiv
Top 0.1%
5.3%
Show abstract

Entamoeba histolytica is a clinically important pathogenic eukaryote and the causative agent of amoebic dysentery. Entamoeba dispar, a nonpathogenic commensal species that resides in the human colon, is the closest sibling species, and serves as an appropriate comparator for genome-wide analysis. Although the genome of E. histolytica is approximately 26.9 Mb, and the largest known genome within the genus, that of E. invadens, is approximately 40.9 Mb, obtaining high-quality assemblies in this genus has remained challenging due to extensive repetitive regions, tRNA gene arrays, and aneuploidy. Here, we used PacBio HiFi sequencing to assemble the genomes of the pathogenic E. histolytica and the nonpathogenic E. dispar. We reconstructed all 36 chromosomes of E. histolytica and 35 chromosomes of E. dispar, assembling each as a single continuous DNA sequence (contig). The two species exhibited high genome-wide nucleotide similarity and conserved synteny at the amino acid level. At one end of each chromosome, we identified tRNA arrays, whereas the opposite end lacked such arrays, resulting in an asymmetric chromosomal architecture. Analysis of unique-read depth revealed widespread aneuploidy in both species: E. histolytica is predominantly tetraploid, whereas E. dispar is diploid, a conclusion further supported by SNP allele-frequency distributions. These assemblies provide a robust foundation for comparative genomics in Entamoeba and offer detailed insights into chromosome-end structure and ploidy.

10
GENE-FAM: An automated pipeline for mining gene families and its application to MADS-box genes in Cannabis sativa

Ryan, L.; Trubanova, N.; Pender, G.; Melzer, R.; Hughes, G. M.; Schilling, S.

2026-06-15 genomics 10.64898/2026.06.10.731441 medRxiv
Top 0.1%
5.2%
Show abstract

Understanding how gene families evolve can offer great insight into adaptation at the phenotypic and ecological levels. This is particularly true in plants, where transcription factor gene families are often targeted for breeding programs to improve the agronomic traits of economically important crops. While recent advances in next generation sequencing have accelerated the wealth of genomics data, there remains a lack of accessible and reproducible genome mining pipelines tailored for gene family characterisation. Here, we address this gap by developing GENE-FAM, an automated, scalable and open-source pipeline designed to mine and predict gene families based on conserved domains and motifs. To illustrate its application, we apply GENE-FAM to annotate MADS-box transcription factor genes across multiple Cannabis sativa genomes. A comprehensive set of MADS-box genes was identified across three C. sativa cultivars, including both previously annotated and newly predicted genes. Through phylogenetic analyses, we confirm that all type II MADS-box gene subfamilies represented in flowering plants are present in C. sativa. Comparing our annotations with those of Arabidopsis thaliana and Solanum lycopersicum revealed that while most MADS type II families are highly conserved, SEPALLATA-like genes have undergone diversification in C. sativa. Together, these results demonstrate the application of GENE-FAM for genome-wide identification and characterisation of gene families in non-model species, revealing novel insights into MADS-box gene family evolution in C. sativa.

11
Valeriana officinalis genome sequence reveals candidate genes for valerenic acid biosynthesis and flavonoid metabolism

de Oliveira, J. A. V. S.; Baez, M.; Pucker, B.

2026-08-21 genomics 10.64898/2026.08.14.744958 medRxiv
Top 0.1%
5.2%
Show abstract

Valeriana officinalis is the scientific name for valerian, a plant known for producing valerenic acid, a compound with anxiolytic properties. Anxiety disorders represent a significant global health crisis, impacting everyday lives. As the global demand for natural, non-synthetic anxiety treatments rises, V. officinalis has emerged as a promising, yet underutilized, medicinal resource. Understanding its genome is the first step toward unraveling the biosynthetic genes underlying valerenic acid production, facilitating further research into its production. Here, we report the first genome sequence of valerian, with an assembly size of 3.3 Gbp and an N50 of 110.8 Mbp, and its corresponding annotation with 96.6% completeness, providing a foundational resource for studying the genetic basis of specialized metabolism in valerian. The value of this genome sequence for discoveries in specialized metabolism is demonstrated by the identification of the flavonoid biosynthesis gene repertoire and the selection of strong candidate genes for valerenic acid biosynthesis. This genome sequence holds the potential to support future functional studies aimed at elucidating the regulation of medically relevant metabolite pathways in V. officinalis.

12
Assembling and annotating the Tainung 67 rice genome to trace its japonica and indica ancestries

Panibe, J. P.; Wang, L.; Wang, T.-Y.; Wang, C.-S.; Lu, M.-Y. J.; Li, W.-H.

2026-07-17 genomics 10.64898/2026.07.15.738793 medRxiv
Top 0.1%
4.8%
Show abstract

Here we reported the 409.0 Mb assembly of the Tainung 67 (TNG67) genome, an early hybrid of japonica and indica cultivars. We used different platforms to sequence the TNG67 genome with a total coverage of 209.8x, from a combination of Illumina paired-end reads, Illumina mate-pair reads, and Oxford Nanopore Technology long reads. The assembly has an N50 of 32.3 Mb with the longest scaffold = 45.2 Mb. We have annotated 38,938 genes where 28,376 (72.9%) finished with the Blast2GO annotation stage. There were 4,821 genes, (98.4%) of the total gene count that have complete BUSCOs. 48.33% (197.5Mb) of the TNG67 genome is composed of repeats. We predicted a total of 527 blast-resistant genes in TNG67. We also analyzed a total of 20 grain size genes and 38 photoperiod-related genes of TNG67, Nipponbare and TN1, and grouped them in terms of whether the TNG67 genes are more similar to Nipponbare, more similar to TN1, hybrid, same, or unique. We also determined whether the sequences of the TNG67 genome were derived from its japonica or indica ancestors. This TNG67 genome may help rice researchers improve yield, develop resistance against biotic and abiotic stress, and understand the evolution of a hybrid cultivar of japonica and indica.

13
A subgenome-resolved and chromosome-scale reference genome assembly of allotetraploid wheat wild relative Aegilops peregrina

Singh, J.; Gudi, S.; Maughan, P. J.; Gill, U.; Gupta, R.

2026-08-30 genomics 10.64898/2026.08.28.747929 medRxiv
Top 0.1%
3.9%
Show abstract

Aegilops peregrina is a wild allotetraploid wheat wild relative and an important source of genetic diversity for stress tolerance and agronomic traits. Here, we report a subgenome-resolved, chromosome-scale reference genome assembly of a drought tolerant and stem rust resistant Ae. peregrina accession PI 604178 generated using PacBio HiFi and Hi-C sequencing. The 10.13 Gb assembly contains 98.81% of sequence anchored to 14 pseudomolecules representing the seven S and seven U chromosomes, with contig and scaffold N50 values of 25.84 and 746.48 Mb, respectively. The assembly achieved a consensus quality value of 74.61, 97.83% k-mers completeness, and 99.9% BUSCO completeness. LTR Assembly Index values of 20.43 and 18.79 for the S and U subgenomes, respectively, further supported high continuity across repeat-rich regions. Repetitive elements comprise 85.93% of chromosome-anchored assembly. We annotated 59,910 high-confidence protein-coding genes, with comparable gene representation across the two subgenomes. This reference genome provides a high-quality genomic framework for comparative analyses, characterization of important loci regulating agronomic and resilience related traits, and sequence-guided exploitation of Ae. peregrina allelic diversity for wheat improvement.

14
The epigenomic landscape of deep lineage divergence: The case of the European sea bass

Longo, A.; Babbucci, M.; Jiao, Z.; Ferraresso, S.; Franch, R.; Bortoletti, M.; Bertotto, D.; Faggion, S.; Ilsley, G. R.; Papadogiannis, V.; Manousaki, T.; Kristoffersen, J.; Tsigenopoulos, C. S.; Macqueen, D. J.; Bargelloni, L.

2026-07-28 evolutionary biology 10.64898/2026.07.27.738391 medRxiv
Top 0.1%
3.9%
Show abstract

BackgroundUnderstanding the role of non-coding genomic variation in speciation remains a major challenge in evolutionary biology. Here, we investigated whether regulatory elements contribute to this process between Atlantic and Mediterranean lineages of European sea bass (Dicentrarchus labrax), a well-characterized case-study near speciation where barriers to introgression exist in the presence of connectivity between diverging populations. ResultsWe generated a novel, highly contiguous genome assembly, which was annotated at the epigenomic level using ATAC-seq and ChIP-seq with six embryonic developmental stages and five tissue types in adult fish, identifying thousands of promoters, enhancers, and open chromatin regions. Integrating this annotation with whole-genome sequence data from 65 individuals across three geographically distinct populations, we identified 57,505 outlier SNPs and 332 structural variants (SVs) showing elevated differentiation between Atlantic and East Mediterranean lineages. Outlier SVs affected key regulatory elements and coding genes, while outlier SNPs were enriched in regulatory elements, particularly enhancers active in adult tissues. Local genomic divergence correlated positively with regulatory element density, especially on chromosomes 1, 9, and 18, which are enriched in genes related to osmoregulation, immune response, and oxidative stress -- processes relevant to adaptation across contrasting marine environments. ConclusionsThese findings support a major role for regulatory variation in driving deep lineage divergence through local adaptation.

15
Annual life-history strategy hitchhikes low-light adaptation in a clonal seagrass

Zhang, X.; Zhang, F.; Suonan, Z.; Zhang, Y.; Li, Y.-L.; Li, X.; Kim, S. H.; Zhou, Y.; Lee, K.-S.; Yu, L.

2026-07-03 evolutionary biology 10.64898/2026.06.29.735426 medRxiv
Top 0.1%
3.9%
Show abstract

While life-history strategies are typically fixed within species, evolutionary transitions between perenniality and annuality can occur. In clonal seagrasses, annual and perennial plants often coexist in the same population, providing a unique model for studying the genetic basis of this transition. Two seagrass Zostera marina populations in South Korea display a striking dichotomy: shallow-water sub-populations follow a typical perennial strategy, whereas their deep-water counterparts are annual. Here we show that this shift from perenniality to annuality, potentially caused by the SAPK7 gene, is genetically coupled with the CAO gene, which is under strong positive selection for low-light adaptation. The up-regulation of the SAPK7 gene triggers early flowering in seedlings, before the formation of any lateral shoots via asexual reproduction. In this special case where the genet contains only one ramet, the post-reproductive death of the ramet is equivalent to the death of the whole genet, which explains the annual phenotype. Our findings reveal a mechanistic example where annuality overcomes perenniality by hitchhiking on a positively selected gene. Given that the ancestral state of plants is perennial, this coupling of annuality with beneficial alleles may represent one of the pathways for the repeated evolution of annual life histories across flowering plants.

16
Whole genome sequencing and variant discovery in 344 global grasspea (Lathyrus sativus L.) lines

Schreiber, M.; Staples, J.; Emmrich, P. M. F.; Edwards, A.; Martin, C.; Bayer, M.; Raubach, S.; Kilian, B.; Shaw, P. D.

2026-06-09 genomics 10.64898/2026.06.05.730453 medRxiv
Top 0.1%
3.3%
Show abstract

The rapid expansion of genomic data resources for major crops is opening new options for crop improvement, while resources for most underutilised crops lag behind, risking a widening gap in crop improvement. One of these underutilised crops is grasspea (Lathyrus sativus), an ancient crop with modern cultivation centred on South Asia and Ethiopia. We conducted whole genome shotgun sequencing on a global collection of 344 grasspea lines, producing over 152 billion reads. Following variant discovery and filtering we created a single nucleotide polymorphism (SNP) marker set of over 1.5 million SNPs. This is a resource of major significance for this crop which can help unlock its breeding potential through marker development and the identification of genes controlling agronomically important traits.

17
Comprehensive transcriptome data of melittin- and un-treated murine hepatoma Hepa 1-6 cells

Zhang, R.;Zhang, Y.;Zang, H.;Lou, J.;Li, Y.;Jiang, J.;Chen, D.;Yan, T.;Guo, R.

2026-06-30 Cancer Biology 10.64898/2026.06.25.734412 medRxiv
Top 0.1%
3.3%
Show abstract

Melittin, the principal bioactive peptide of bee venom, exerts potent antitumor activity against hepatocellular carcinoma (HCC). However, the comprehensive transcriptomic alterations it elicits in hepatoma cells remain poorly characterized. Here, we present an integrated transcriptome dataset from melittin- and un-treated murine Hepa 1-6 hepatoma cells, encompassing messenger RNA (mRNA) and microRNA (miRNA) expression profiles. Cells were exposed to 4 g/mL melittin in serum-free DMEM for 20 min, and total RNA was subjected to ribosomal RNA-depleted strand-specific RNA sequencing on an Illumina NovaSeq6000 platform (paired-end 150 bp) and small RNA sequencing on an Illumina HiSeq2500 platform (single-end 50 bp). Raw data were processed using Cutadapt to remove adapters and low-quality reads, yielding clean datasets with Q20 [≥] 99.85%, Q30 [≥] 98.48%, and valid data ratios exceeding 85%. All raw and processed sequencing data are publicly available. This transcriptomic resource provides a valuable resource and basis for elucidating the regulatory networks underlying melittin-induced anti-hepatoma effects. DatasetThe dataset can be accessed through the National Genomics Data Center, China National Center website by searching with the BioProject accession number PRJCA065485 Reviewers may use this link for anonymous access during the review process. Direct URL to data: Genome Sequence Archive-CNCB-NGDC. Dataset LicenseCC BY 4.0

18
Whole genome sequences and annotations of Japanese and French strains of Heterosigma akashiwo

Kondo, T.; Sakamoto, M.; Tokumaru, M.; Tanizawa, Y.; Nakamura, Y.; Toyoda, A.; Ueki, S.

2026-08-23 genomics 10.64898/2026.08.19.745619 medRxiv
Top 0.1%
2.7%
Show abstract

High-quality reference genomes provide an essential foundation for elucidating the molecular basis of organismal ecophysiology. Here, we sequenced and assembled chromosome-scale genomes of two Heterosigma akashiwo strains isolated from coastal waters of Japan and France. The assembly sizes were 1.18 Gb and 1.43 Gb for the Japanese and French strains, respectively. The scaffold N50 of the Japanese strain assembly was 66 Mb, whereas the one of the unscaffolded French strain assembly was 33 Mb. To our knowledge, these assemblies represent among the largest and most contiguous genome resources currently available for members of the Stramenopiles (Ochrophyta). Evidence-based gene prediction in the Japanese strain recovered approximately 90% of conserved stramenopile core genes, indicating a highly complete gene repertoire, and was complemented by extensive functional annotation. In the French strain, homology-based gene prediction recovered approximately 80% of conserved core genes. Comparative genome analysis revealed extensive synteny conservation between the two strains, although several putative duplication and translocation events were detected. These genomic resources provide a robust framework for investigating the molecular, cellular, and ecological mechanisms underlying the physiology, adaptation, and bloom-forming capacity of H. akashiwo.

19
Three annotated tiger beetle genomes (Coleoptera, Adephaga, Cicindelidae)

Ramirez, J.; Chou, M.-H.; Gustafson, G.

2026-07-21 genomics 10.64898/2026.07.16.739005 medRxiv
Top 0.1%
2.6%
Show abstract

Advances and accessibility to next-generation technology provide opportunities to sequence non-model organisms. Despite this increase in whole-genome sequencing data, annotation and the production of reference genomes remain limited. Reference genomes are a critical tool for a variety of studies in evolutionary biology, functional genomics, and conservation genetics. Tiger beetles (Cicindelidae) are a diverse and globally distributed family of beetles that serve as bioindicator taxa and flagship species for insect conservation. Here, we report highly complete, contiguous, and annotated genome assemblies representing draft reference genomes for three species of tiger beetle spanning the phylogeny. These draft reference genomes are for Audouins night-stalking tiger beetle, Omus audouini; the montane giant tiger beetle, Amblycheila baroni; and the western red-bellied tiger beetle, Cicindelidia sedecimpunctata. Article SummaryTiger beetles are a charismatic group with [~]3000 species distributed globally. Despite their popularity among insect enthusiasts and their role as bioindicators of ecosystem health, the group currently lacks a reference genome. This article outlines genome assembly and annotation for three tiger beetle species that span evolutionary relationships within the lineage. Quality control analyses show that the assemblies are reference quality and demonstrate high contiguity, completeness, and accuracy. The resulting draft genome annotations will be a valuable resource for scientific endeavors and allow for continued research on tiger beetles and their allies.

20
Long-read sequencing based genomic data of a dipluran species, Occasjapyx japonicus

Asano, T.; Toyoda, A.; Hashimoto, K.; Yokoi, K.

2026-07-21 genomics 10.64898/2026.07.16.716246 medRxiv
Top 0.1%
2.5%
Show abstract

We present the genome dataset of a dipluran species, Occasjapyx japonicus, representing the first dipluran genome assembled using HiFi long-read sequencing technology. The assembled genome is approximately 439.3 Mbp in size, comparable to those of other dipluran species available in public databases. The N50 value of 15.5 Mbp exceeds that reported for other dipluran species. The assembled gene set contains 19,635 genes, a number not significantly different from those estimated in previous analyses of two other dipluran species. Functional gene annotation was conducted using predicted amino acid sequences derived from the gene set. BUSCO analysis indicated that the assembled genome contains the majority of conserved core genes. These findings suggest that the O. japonicus genome and associated data are of sufficient quality to serve as a reference genome. The dataset will be valuable for studies in comparative or evolutionary biology, particularly in understanding hexapod evolution and the emergence of insects.